Papers with visual generation

2 papers
Colorism in Multimodal AI: An Empirical Exploration of Socioeconomic Linguistic Bias in Text-to-Image Generation (2026.eacl-srw)

Copied to clipboard

Challenge: Socioeconomic inequalities worldwide are deeply linked to ethnoracial hierarchies and stereotypes, argues a new study.
Approach: They use a Monk Skin Tone scale to benchmark VLMs and annotators . they then use linguistic cues to vary skin-tone representations in text-to-image generation .
Outcome: The study compares 3 small VLMs and 60 human annotators on the monk skin tone scale with 210 occupations and produces over 2,500 portraits across 3 large VLM models.
UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation (2024.emnlp-main)

Copied to clipboard

Challenge: e-commerce tasks such as multimodal retrieval and multimodal generation are largely ignored due to the diversity of the multimodal fashion domain.
Approach: They propose a framework that integrates image generation with retrieval and text generation tasks.
Outcome: The proposed framework outperforms state-of-the-art models across fashion tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations